Ell [Mon, 18 May 2020 06:53:58 +0000 (09:53 +0300)]
HSV, HSL, HCY: wrap hue around during conversion to RGB
In the conversions from HSV, HSL, and HCY to RGB, wrap the hue
value around to the [0,1) range, instead of producing unspecified
results outside this range. In particular, hue=1.0 may arise when
going through lower precision, such as when decomposing/recomposing
an 8-bit image in GIMP (see gimp#5097).
Øyvind Kolås [Fri, 15 May 2020 00:56:46 +0000 (02:56 +0200)]
babl: adjust search depth for fishes in first pass
Search first only to a depth of 2, and accept the fastest found if any
faster than reference is found, if not search to depth 5. We lose out in
the unlikely to be frequent cases where 2-step babl fishes exist and a 3
step fish is faster. And gain the fishes where 4 step is faster than 3
(which was the old first pass search depth).
Øyvind Kolås [Thu, 14 May 2020 15:53:50 +0000 (17:53 +0200)]
babl: slightly pad source buffers for creating conversions
This is done because many babls conversions get optimized by C compilers
to read 16bytes of data at a time. Causing valgrind to report
"Invalid read of size 16" as a false positive. By padding the data at
least when creating conversions we mask the false positives.
avx2-int8: add gamma u8 -> linear float conversions
Add AVX2 conversions from u8 Y', Y'A, R'G'B, and R'G'B'A to float
Y, YA, RGB, and RGBA, respectively. The conversions use an LUT
together with the AVX2 gather instructions to process 8 values a
once. Depending on the formats and cache utilization, the new
conversions are between 1.25x to 2.2x faster than the existing
conversions.
Øyvind Kolås [Sat, 15 Feb 2020 06:22:25 +0000 (07:22 +0100)]
babl: babl-fish-path improve portability for ppc
Quoting his explaination for why this fixes things (the other half of
the same function was already portable.)
"This breaks on PowerPC 32-bit, because the calling convention for
passing a union is to pass a pointer to a temporary copy of the union.
This pointer isn't as a function pointer. On some other platforms, the
call would just copy the union, which is like copying the function
pointer. On all platforms, the compiler doesn't check the type, because
babl_conversion_new() has a va_arg(3) prototype."
Portabilitiy issue figured out by George Koehler, This fixes issue #24.
Øyvind Kolås [Sun, 12 Jan 2020 22:47:08 +0000 (23:47 +0100)]
meson: globally opt out of unsafe math optimizations
Thus closing issue #49, with a decision to opt for fully predictable
math, and increasing the ability to extend the use of hashes in result
image tests in GEGL.
Øyvind Kolås [Fri, 8 Nov 2019 18:07:47 +0000 (19:07 +0100)]
build: opt out of unsafe math optimizations in reference and base
This makes the reference code paths used for verifying conversions in
extensions not involve for instance fast reciprocal approximations, see issue
#49. The extensions are still compiled with full optimizations.
Øyvind Kolås [Fri, 8 Nov 2019 12:54:15 +0000 (13:54 +0100)]
babl: adjust default BABL_TOLERANCE
This as a start of fixing issue #49, lower precision code gets generated
on AMD EPYC, due to use of rcpps to get a reciprocal - which has lower
precision on AMD EPYC. The conversions with AMD EPYC gets included with
a small margin - so we should not be shedding many other fast conversions
due to this.
Øyvind Kolås [Tue, 20 Aug 2019 00:54:51 +0000 (02:54 +0200)]
babl: register gray and gray alpha shortcuts between spaces
The formats "Y float", "YA float" and "YaA float" between different RGB
and gray spaces are all exactly the same. This registers memcpy shortcuts
for these combinations for all spaces. This avoids temporarily expand
into RGBA for many babl path fishes.
Ell [Mon, 19 Aug 2019 14:33:29 +0000 (17:33 +0300)]
avx2, two-table: don't segfault for NaN input
In the AVX2 and two-table linear-float => gamma-int8 conversions,
tweak the input bounds-check to handle NaN values. NaN would
previously lead to an out-of-bounds table lookup, and a segfault
(see issue #43).
Øyvind Kolås [Sun, 18 Aug 2019 23:00:11 +0000 (01:00 +0200)]
babl: reduce default max path length to 3 (5)
With the space invasion progressing, we're starting to have
many more conversions to search through - in some cases
looking for conversions - when there are no conversions
become too expensive with an exhaustive search of all
conversions up to 7 steps. (it is likely possible to reduce
the amount of conversions tried).
As well re-enable handling of grayscale ICC profiles - since
the formerly patological tests on gimp now pass.
Øyvind Kolås [Sun, 18 Aug 2019 22:01:13 +0000 (00:01 +0200)]
babl: add support for grayscale spaces
A grayscale space is just like an RGB space, but has a default set
of chromaticities for R,G,B. When ICC profiles are loaded only the TRC
is considered is considered (we could also include the whitepoint),
blackpoint tag is ignored (and should be baked into the used ICC
profile instead.)
This works with GIMP-2.10 but master of GIMP currently hangs when
trying to load a grayscale jpeg with attached grayscale ICC profile,
The if #0 on line 999 of babl/babl-icc.c needs to be turned into a
1 to enable grayscale icc profiles for further testing.
Øyvind Kolås [Fri, 16 Aug 2019 21:45:18 +0000 (23:45 +0200)]
babl: simplify logic in babl_epsilon_for_zero
We do not need to treate positive/negative avoided infinities differently,
the result is the same as long as we are consistent, and only using the
positive epsilon value leads to slightly simpler per-pixel conversion code
in both directions.
Øyvind Kolås [Tue, 6 Aug 2019 11:46:35 +0000 (13:46 +0200)]
meson: stop overriding libdir
Let meson do it's defaults, this mean we get installed in
prefix/lib/x86_64-linux-gnu/ instead of prefix/lib/
apparently which apparently is better - but different from
what autotools used to do; and leads to other paths also
needing potential adjustment.
Øyvind Kolås [Fri, 2 Aug 2019 11:53:44 +0000 (13:53 +0200)]
build: issue #41 builds without mmx/sse etc extensions do not work
The warnings about the SIMD flags specific variables not being set goes
away with this commit. Setting empty strings do not work since then gcc
ends up looking for input files that are the empty string, thus this patch
re-adds -Wall instead of setting an empty string.
This silences a false warning from gcc, where it doesn't recognize the type
of control flow occuring with the pre/post switches in the reference
conversion.
Add AVX2 conversions from Y float, YA float, RGB float, and RGBA
float, to Y' u8, Y'A u8, R'G'B' u8, and R'G'B'A u8, respectively.
The conversions use a lookup table, similarly to the two-table
conversions, indexed using AVX2's 256-bit gather instruction, which
allow us to process 8 floats at once. Over here, this conversion
is ~5x faster than the SSE conversions.
Note, however, that unlike two-table, we don't use a second
"feedback" table to correct the result, leading to an off-by-one
conversion error with a probability of ~0.1%, using a 2^16-element
table. This error rate is low enough for babl to use the
conversion, but might still be a bit too high regardless; it can be
further reduced with a bigger table.